Papers by Zubin Trivadi Aysola
Rejected Dialects: Biases Against African American Language in Reward Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Preference alignment via reward models can introduce new biases, hindering reward models’ fairness and equity. |
| Approach: | They propose a framework for evaluating dialect biases in reward models and conduct a case study on biase . they compare reward models' preferences and behavior on paired White Mainstream English and machine-translated and human-written AAL corpora. |
| Outcome: | The proposed framework evaluates dialect biases in reward models and compares them with paired White Mainstream English (WME) and machine-translated and human-written AAL corpora. |